Papers with compositional understanding
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage (2026.acl-long)
Copied to clipboard
| Challenge: | Vision-Language Models excel at photorealistic generation, but struggle to represent abstract meanings. |
| Approach: | They propose a benchmark that replaces high-fidelity visual detail with schematic iconicity by generating paired, sense-anchored visualizations for literal and idiomatic readings. |
| Outcome: | The proposed benchmark replaces high-fidelity visual detail with schematic iconicity by generating paired, sense-anchored visualizations for literal and idiomatic readings. |
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing fine-tuning approaches for compositional understanding compromise performance in zero-shot multi-modal tasks. |
| Approach: | They propose a method to enhance compositional understanding in pre-trained vision and language models without sacrificing performance in zero-shot multi-modal tasks. |
| Outcome: | The proposed method achieves compositionality on par with state-of-the-art models and retains strong multi-modal capabilities. |
Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality (2026.acl-long)
Copied to clipboard
| Challenge: | a contrastive learning approach for vision-language models is needed to capture compositional information. |
| Approach: | They propose a framework that masks compositional concepts in one modality and reconstructs them conditioned on full contextual information from the other . |
| Outcome: | The proposed framework enhances compositionality in visual language models and improves their ability to capture syntactic structure and linguistic information. |